Papers with continuous reward modeling

1 papers
ARF-RLHF: Adaptive Reward-Following for RLHF through Emotion-Driven Self-Supervision and Trace-Biased Dynamic Optimization (2026.acl-long)

Copied to clipboard

Challenge: prevailing RLHF methods such as PPO and DPO depend on large-scale binary preference annotations.
Approach: They propose a method which converts natural feedback into continuous preference trajectories and optimizes them using the novel TraceBias algorithm.
Outcome: The proposed approach outperforms PPO and DPO in a variety of domains and improves alignment by up to 7.6% across diverse LLMs and preference domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations